NOTE

1.12 Elasticsearch refresh

English translation of the original VNote ‘Elasticsearch refresh’, preserving its structure and historical learning notes.

Elasticsearch / SearchCreated Updated 1 min readhistorical

This is a historical learning note and may contain outdated or incomplete understanding.

1. What Is refresh?

  • An Elasticsearch operation that writes data from the memory buffer into the filesystem cache, after which it becomes searchable.

2. Why Is refresh Needed?

  • Make the index searchable.
    • Note that after refresh, the data is in the filesystem cache and has not been persisted to disk, so data may be lost. Therefore Elasticsearch provides Elasticsearch translog.
    • Only after Elasticsearch flush is performed will segment files be persisted to disk.

3. Why Not fsync Directly to Disk?

  • The refresh operation is lighter than fsync, although it sacrifices data reliability.

4. When refresh Is Triggered

4.1. Manual

  • POST /my_index/_refresh can be used to force a refresh.

4.2. Automatic

  • By default, refresh happens once per second.
  • The refresh interval can be configured.
    PUT /my_index/_settings
        {
          "index" : {
            "refresh_interval" : -1
          }
        }

5. refresh Process

Write data from the memory buffer into the filesystem cache. A Lucene segment file is generated at this point.

6. References

Discussion

Sign in with GitHub to comment. Discussions are stored as GitHub Issues.View on GitHub